Skip to content

feat(webapp): dashboard agent — Watch - #4525

Open
kathiekiwi wants to merge 109 commits into
feat/dashboard-agent-uifrom
feat/dashboard-agent-flows-watch
Open

feat(webapp): dashboard agent — Watch#4525
kathiekiwi wants to merge 109 commits into
feat/dashboard-agent-uifrom
feat/dashboard-agent-flows-watch

Conversation

@kathiekiwi

@kathiekiwi kathiekiwi commented Aug 6, 2026

Copy link
Copy Markdown
Collaborator

Watch is the agent noticing something later: you ask it to tell you when a condition holds, and it answers when it does — or when it can't any more.

A watch is a durable one-shot promise. The condition is checked on a schedule by deterministic code (no LLM in the checks), the answer lands in the chat once, and then the watch is over. Ten kinds: three on a run, five on a queue, error recurrence, health recovery.

Stack

Stacked on #4529 (UI), which is stacked on #4418 (chat, reports, investigate). Merge those first. #4516 (storybook gallery) sits on top of this branch.

How to review

GUIDEBOOK.md on this branch is the behaviour reference — it states the conditions rather than the code, so you can predict what happens without running anything. "The ten watch kinds, and what makes each fire" and "Creating a watch" describe exactly this PR, and the tables there are the spec the code is written against.

What's inside

  • Ten watch kinds, one deterministic check each (dashboardAgentWatch*Checks.ts), with the spec union in dashboard-agent-contracts/src/watch.ts.
  • Scheduling — each watch schedules its own next check; due watches of one (environment, cadence) group can be checked together in one batch pass, with a sweep as the backstop for expiry, redelivery and retention.
  • Delivery — the in-chat wake and card, an optional email alert (new DASHBOARD_AGENT_WATCH alert channel, so it shows on the project's Alerts page with one-click unsubscribe), and an optional investigation when the outcome needs attention.
  • Submission ledgerwatch_submissions, keyed (chat_id, client_request_id), so a retried card submission replays the recorded outcome instead of creating a second watch.
  • Watch token — a delayed-execution credential accepted only by the watch endpoints, re-checked against the user's live access on every tick.
  • Unread work — the panel polls for wakes that landed while it was closed, so a chat can go unread and light the launcher dot.

Key decisions

A check result is a 4-way, and only two of them are verdicts. satisfied / terminal_unsatisfied are answers; pending and unavailable are not. Any exception inside any check is caught in one place and becomes unavailable with an unverified observation — a check that failed is never evidence.

A completed window is an answer, and whether it is good or bad news is declared per kind, never inferred. There is a table for that in the guidebook: run_failed completing its window is good news ("hasn't failed"), backlog_drain completing it is not. One rule overrides the table: a window that completed on an unverified observation is neutral and says only that the watch ended without a confirmed answer. An unreadable source is never a negative answer — and, because investigations only open on attention, it never starts one either.

Identity is (chat, project, environment) plus the condition, enforced by a partial unique index over active rows (watches_chat_active_identity_key), not by the read-then-insert check. Cadence, window, note and ticks are deliberately not part of it. Two different chats may watch the same thing — a watch is a promise to a chat.

The server resolves the target's name, whatever the model calls it. The model can't tell a task queue (task/<id>) from a custom queue, so both spellings are tried and the stored one wins — and the rewrite happens before identity and before the row is written, so the identity, the checks, the link and the wording all see one spelling.

Freshness fences. Depth falls back from the live counter to the newest 60 s ClickHouse bucket, which only counts as current within 60 s of now. A non-current reading at or below the quiet line is refused as unavailable rather than believed, so a stale empty bucket is never read as "drained". The stall streak is the one piece of carried state: it lives in the previous check's facts and freezes on an unreadable reading rather than breaking.

Chain reliability. There is no shared cron — each watch (or batch group) schedules its own next tick, so the failure mode to review is the chain dying. A failed batch check is caught, the next tick is scheduled anyway and the run resolves rather than failing, so the chain survives a check that couldn't run; the sweep re-arms groups and finalizes anything still active past its deadline, even when delivery isn't configured. Wake redelivery is id-deduped rather than conditional, because the sweep can't know whether the user was already told. Access is re-authorized on every check against the primary — replica lag would extend access the user has already lost.

Wording lives in one place. watch-wording.ts is read by the card, banner, toast, email and the agent's own narration, and the numbers come from the frozen observation rather than a fresh read, so a retry produces the same sentence. Replay reproduces the recorded decision instead of deciding again — the transcript is append-once, so a second decision would contradict it forever.

Cancellation is the ending without an answer — no resolution, no wake. One exception, decided during testing: a watch the user cancelled leaves a single neutral transcript line ("Stopped watching …"), keyed off the watch id so a retry can't repeat it. The other four reasons stay silent.

Email is opt-in and only a fired watch emails. An expiry is narrated in the chat and nowhere else. Both gates (agent access, a configured email transport) are checked at subscribe time and again at delivery, and the subscription outcome is frozen on the ledger row so a retry replays it. Neither gate is a plan check.

One watch offer per turn. The prompt and the renderer guard this independently — if the turn already proposed a watch card, the action button is dropped, because the card is the better affordance. Two eval cases pin the prompt side: exactly one offer with the line last and the button after it, and zero offers when the rendered card already carries one — deterministic assertions, over a real-model run.

Testing

Unit tests (vitest, testcontainers, no mocks) under apps/webapp/test/dashboardAgentWatch*.test.ts and internal-packages/dashboard-agent/src/watch-*.test.ts cover the invariants above: the 4-way check results and the freshness fences, identity/dedup and the submission ledger, queue-name resolution, the batch chain surviving a failed check, sweep boundaries and alert-once, tenancy and the watch token's scope, and the wording snapshot. The load-bearing ones were verified by control-breaking the guard first and checking the test goes red.

Live-tested end to end against a local stack, following the guidebook: all ten watch kinds firing and expiring, cancellation, the email pair (a fired watch mails, an expired one does not), and watch recovery from a health report.

@changeset-bot

changeset-bot Bot commented Aug 6, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: f6c90ec

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 27 packages
Name Type
@trigger.dev/core Patch
@trigger.dev/build Patch
trigger.dev Patch
@trigger.dev/python Patch
@trigger.dev/redis-worker Patch
@trigger.dev/schema-to-json Patch
@trigger.dev/sdk Patch
@internal/cache Patch
@internal/clickhouse Patch
@internal/llm-model-catalog Patch
@internal/metrics-pipeline Patch
@trigger.dev/rbac Patch
@internal/redis Patch
@internal/replication Patch
@internal/run-engine Patch
@internal/run-store Patch
@internal/schedule-engine Patch
@trigger.dev/sso Patch
@internal/testcontainers Patch
@internal/tracing Patch
@internal/tsql Patch
@internal/dashboard-agent Patch
@internal/sdk-compat-tests Patch
@trigger.dev/react-hooks Patch
@trigger.dev/rsc Patch
@trigger.dev/database Patch
@trigger.dev/otlp-importer Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@coderabbitai

coderabbitai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: ea43b923-0e58-4423-83a1-821de5df924f

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Walkthrough

Added watch creation and configuration for runs, queues, errors, and health reports. Added scheduled checks, lifecycle handling, wake notifications, automatic investigations, unread tracking, and cross-browser activity polling. Added email, Slack, and webhook alert delivery with subscription management. Added dashboard-agent tools, APIs, persistence, worker tasks, scenario tooling, documentation, and extensive unit and integration coverage.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 64.71% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the dashboard agent Watch feature and matches the primary change.
Description check ✅ Passed The description clearly explains the feature, design decisions, testing, and review guidance, but it omits the template checklist, issue closure, and screenshots sections.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/dashboard-agent-flows-watch

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@kathiekiwi
kathiekiwi marked this pull request as ready for review August 7, 2026 10:20
devin-ai-integration[bot]

This comment was marked as resolved.

coderabbitai[bot]

This comment was marked as resolved.

@kathiekiwi
kathiekiwi changed the base branch from feat/dashboard-agent-flows to feat/dashboard-agent-ui August 7, 2026 12:43
@pkg-pr-new

pkg-pr-new Bot commented Aug 7, 2026

Copy link
Copy Markdown

Open in StackBlitz

@trigger.dev/build

npm i https://pkg.pr.new/@trigger.dev/build@ea98c39

trigger.dev

npm i https://pkg.pr.new/trigger.dev@ea98c39

@trigger.dev/core

npm i https://pkg.pr.new/@trigger.dev/core@ea98c39

@trigger.dev/python

npm i https://pkg.pr.new/@trigger.dev/python@ea98c39

@trigger.dev/react-hooks

npm i https://pkg.pr.new/@trigger.dev/react-hooks@ea98c39

@trigger.dev/redis-worker

npm i https://pkg.pr.new/@trigger.dev/redis-worker@ea98c39

@trigger.dev/rsc

npm i https://pkg.pr.new/@trigger.dev/rsc@ea98c39

@trigger.dev/schema-to-json

npm i https://pkg.pr.new/@trigger.dev/schema-to-json@ea98c39

@trigger.dev/sdk

npm i https://pkg.pr.new/@trigger.dev/sdk@ea98c39

commit: ea98c39

@github-actions

github-actions Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

Observability map

As of f6c90ec.

20/100 over 424 measured of 440 entry points (base 19, up 1)

What this PR changed

route base head now failing
/api/v1/dashboard-agent/watches/:watchId/check (suppressed: error-classification) new 50

FIX FIRST

  • /api/v1/projects/:projectRef/envvars (sensitive) - auth-boundary, request-context
  • /auth/sso (sensitive) - auth-boundary, request-context
  • /_app/orgs/:organizationSlug/settings/team (sensitive) - error-classification, auth-scope, request-context

AUDIT 3 of 50 sensitive mutations record an actor. 47 without one.
CONTEXT 21 of 424 entry points name a tenant on a failure path. 324 appear only here, 39 of them sensitive, in the JSON rather than the fix list.

What the score is made of
CHECKS
  error-classification  177 applicable, 101 pass,   0 sole, global without it 12
  auth-boundary          62 applicable,  57 pass,   0 sole, global without it 17
  auth-scope             19 applicable,  17 pass,   0 sole, global without it 20
  request-context       424 applicable,  21 pass, 225 sole, global without it 64
  audit-trail            50 applicable,   3 pass,   0 sole, not in the score

The score and findings here are report-only and never gate the merge. Separately, a required test suite keeps this tool's symbol and route lists in sync with the code they name, and can fail a pull request that renames or removes a symbol they reference, or that adds the first route with a segment they anticipate. Each failure names the list to edit. The rules and their reasons: internal-packages/observability-map/README.md.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

@kathiekiwi
kathiekiwi force-pushed the feat/dashboard-agent-flows-watch branch from 712396b to e110e90 Compare August 8, 2026 12:05
@kathiekiwi
kathiekiwi force-pushed the feat/dashboard-agent-ui branch from bd4d4a0 to 887f5b6 Compare August 8, 2026 12:05
@kathiekiwi
kathiekiwi force-pushed the feat/dashboard-agent-flows-watch branch from c0f0058 to e7432a8 Compare August 8, 2026 14:30
@kathiekiwi
kathiekiwi force-pushed the feat/dashboard-agent-ui branch from 887f5b6 to 17a0f07 Compare August 8, 2026 14:30
@coderabbitai

coderabbitai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor

Note

GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer.

coderabbitai[bot]

This comment was marked as resolved.

…s-watch

# Conflicts:
#	apps/webapp/app/components/dashboard-agent/DashboardAgentPanel.tsx
The offer line is the last sentence, the actions block comes after it. Also widens the no-duplicate-offer exemption to a health report card that already offers a recovery watch.
Keep the precomputed blocksByPart (winners filter already applied) and
pass watchOfferedInTurn through the new slot factory; actions render
last inside the wake slot too.
…s-watch

# Conflicts:
#	internal-packages/dashboard-agent/src/__snapshots__/prompt-prefix.test.ts.snap

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 new potential issue.

Open in Devin Review

Comment on lines +853 to +860
} finally {
// No `onTurnComplete` on an action, so the guard runs here: a card left at
// in_progress is a spinner nothing else will ever stop. `closeCard` is the settle,
// so nothing settles the row separately first.
await closeCard(investigationId, answered ? [...uiMessages, answered] : uiMessages);
// Only once the close committed: a retry needs the tracked entry to still be there.
clearOpenInvestigations(chatId);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 A transient model failure still settles the consented investigation's card before the action retries

In conductWatchInvestigation the closing write runs in a finally, so any throw out of streamText / chat.pipe (rate limit, provider blip, network) settles the row and appends the closing card before the error propagates and the task retries.

On the retry, the dedupe guard reads uiMessages (the session snapshot), which will not contain the closing card that was written straight to the read-model DB — settleInvestigationCard appends there, not to the session's out stream. So alreadyAnswered is false and latestCards(uiMessages) has no entry for the id, meaning the retry re-runs the whole investigation, revises the already-terminal row via render_view, and calls closeCardInTranscript again (which bumps the revision even though appendChatMessageOnce dedupes the message). The user can see the card go inconclusive and then move again, and the row's revision can end up ahead of the card the transcript holds.

The tests cover the case where the closing card itself can't be written (internal-packages/dashboard-agent/src/watch-actions.test.ts — "fails the action when the closing card can't be written"), but not the case where the model call throws first. Worth deciding whether the settle should be skipped when the turn is going to be retried, e.g. only closing on a clean exit and leaving the stale-investigation sweep to handle a crashed run.

Open in Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant